Back

Nature Biotechnology

Springer Science and Business Media LLC

Preprints posted in the last 90 days, ranked by how well they match Nature Biotechnology's content profile, based on 172 papers previously published here. The average preprint has a 0.17% match score for this journal, so anything above that is already an above-average fit.

1
Reference-free protein sequencing by consensus assembly of redundant de novo peptide reads

Nilsson, A.; Sporre, E.; Schulte, D.; Snijder, J.; Edfors, F.; Käll, L.

2026-08-19 bioinformatics 10.64898/2026.08.13.744110 medRxiv
Top 0.1%
34.1%
Show abstract

Reading a proteins sequence from tandem mass spectra without a reference is limited by single-spectrum accuracy, most acutely across the hypervariable complementaritydetermining regions of antibodies. Broadly specific proteases tile a protein with long, overlapping peptides, so every residue is covered by many independent de novo reads. borgonovo assembles their per-step probability profiles into a reference-free per-residue consensus, seeding templates from mass-closure-consistent reads and recruiting the rest by substitution- tolerant alignment and per-column voting. Re-decoding each spectrum with a prior from its consensus position lifts amino acid accuracy on placed spectra from 0.80 to 0.87. On the therapeutic antibody trastuzumab, nine proteases cover its heavy and light chains completely at 0.88 fixed-window identity, and 0.93 on the pruned assembly once local indels are accommodated. Applied unchanged to five secretome proteins and trastuzumab with three proteases, it reaches 0.87 mean fixed-window identity over 82% coverage. borgonovo is open source and works with most de novo sequencers, so redundant digestion turns any of them into a protein sequencer where no reference exists.

2
Design and Assembly of Combinatorial DNA Barcodes for Probe-based Genomics Applications

Goode, Z.; Tiedemann, E.; Ben Ameur, L.; Pavan, K.; Young, K.; Sek, M.; Nevue, A.; Zhu, J.; Houghton, J.; Fu, Y.; Boisvert, H.; Saunders, A.

2026-08-26 genomics 10.64898/2026.08.22.746475 medRxiv
Top 0.1%
26.7%
Show abstract

Probe-based genomics technologies are extending molecular analysis into intact tissues and fixed cells, yet strategies to decode complex experimental conditions encoded in cellular RNA remain limited. Here we present a modular framework that integrates custom software tools with purpose-built cloning reagents to design, assemble, validate, and deploy combinatorial DNA barcodes. Combinatorial barcodes comprise spatially adjacent collections of known sequences, enabling millions of unique molecules to be efficiently distinguished using a limited set of probes. Our software tools integrate with optimized assembly plasmids and whole plasmid long-read sequencing for high-fidelity construction and structural validation of diverse combinatorial barcode architectures. Assembled barcode libraries are flexibly transferred into user-modified expression vectors to support diverse downstream experimental applications. We showcase the versatility of this framework by assembling two structurally distinct combinatorial barcode libraries, each containing millions of unique sequences. Following rabies virus-based delivery to the mouse brain, we validate in vivo decoding of a combinatorial barcode architecture capable of distinguishing ~16.3 million expressed RNAs through probe-based in situ sequencing. Our framework for flexible and accurate combinatorial barcode construction fills a technically demanding niche delivering cost-effective molecular reagents for multiplexed experimentation on current and evolving probe-based genomics platforms.

3
Direct comparison of CRISPR knockout and interference with Perturb-seq

Drepanos, L. M.; Escude Velasco, B.; Chase, A.; Srikanth, S.; Gatzen, M.; Rickner, H. D.; Dubinsky, D.; Navia, A. W.; Winter, P. S.; Shibue, T.; Yates, K. B.; Doench, J. G.

2026-07-04 genomics 10.64898/2026.07.04.736492 medRxiv
Top 0.1%
22.2%
Show abstract

CRISPR knockout (CRISPRko) and CRISPR interference (CRISPRi) are two workhorse technologies for loss-of-function studies, yet direct comparisons between the two are scant relative to their widespread adoption. Here, we establish benchmarking libraries for Cas9-based CRISPRko and CRISPRi screens using Perturb-seq as the read-out. For both modalities, we observe consistent transcriptional signatures among cells with the same genes perturbed, strong evidence of on-target signal. We also examine tradeoffs between modalities: while CRISPRi guides demonstrate heightened rates of off-target activity, we also observe artifacts stemming from the cellular response to double-stranded breaks with the use of CRISPRko. The libraries and analyses presented here will be a useful benchmarking and de-risking resource for any group preparing for a large-scale Perturb-seq screen.

4
Spatial microRNA profiling at single-cell resolution by in situ barcoded extension

Robles-Remacho, A.; Zou, Y.; Jensen, A.; Tricopoulos, C.; Grillo, M.; Nilsson, M.

2026-08-20 genomics 10.64898/2026.08.12.744364 medRxiv
Top 0.1%
18.5%
Show abstract

The spatial organization of post-transcriptional regulation is a fundamental yet difficult to access layer of tissue biology. MicroRNAs (miRNAs) are small RNAs with a key role in post-transcriptional regulation, but their short length has excluded them from spatial profiling technologies, leaving them largely unexplored in spatial transcriptomics. Here, we introduce miR-Space, a method that converts individual miRNAs into extended, uniquely barcoded molecules directly in tissue, enabling their spatial detection by in situ sequencing. Across 30 mouse and human brain sections, miR-Space enabled highly multiplexed miRNA profiling at single-molecule and single-cell resolution, joint analysis with mRNA, and implementation on the automated Xenium platform. miR-Space resolved major anatomical regions and cell populations from spatial miRNA expression, identified reproducible cell-associated miRNA signatures, and uncovered previously unknown spatial and cellular distributions of multiple miRNAs. Together, these capabilities establish miR-Space as a framework for integrating miRNAs into spatial transcriptomics, enabling spatial miRNomics at anatomical and single-cell resolution.

5
Analog intrinsic recoding measures RNA dynamics without chemical conversion

Hansen, L. N.; Couvillion, M. T.; McShane, E.; Azbukina, N.; Treutlein, B.; Churchman, L. S.

2026-07-28 genomics 10.64898/2026.07.27.741041 medRxiv
Top 0.1%
16.9%
Show abstract

Steady-state RNA abundance measurements mask the synthesis and decay rates that shape gene expression. Analog intrinsic recoding sequencing (AIR-seq) repurposes the base-pairing properties of N4-hydroxycytidine (NHC) to mark newly synthesized RNA with C-to-T and T-to-C mismatches in standard RNA-seq libraries, eliminating chemical conversion and enrichment common to RNA metabolic labeling methods. NHC-AIR-seq resolves bulk and single-cell RNA dynamics, improves RNA velocity inference, reveals hidden regulation, and brings RNA kinetics into routine transcriptomic workflows.

6
In situ Discovery of Immune Repertoire Reveals Antitumor Immunity and Therapeutic Antibodies

Zhang, H.; Wang, P.; Zhao, Y.; Yang, L.; Xue, T.; Liu, L.; Zhao, Y.; Zhang, Z.; Ma, J.; Zeng, B.; Zhang, P.; Wang, C.; Pan, D.; Gao, Z.; Liu, Z.; Zeng, Z.

2026-08-19 bioinformatics 10.64898/2026.08.11.744176 medRxiv
Top 0.1%
16.8%
Show abstract

Spatial transcriptomics offers a glimpse into the immunology of tissues. However, limitations in spatial transcriptomics preclude the detection of highly diverse, low-abundance, and previously unknown sequences, including immune repertoires and microbiota. Here, we introduce Archimap, a spatial transcriptomic platform that simultaneously profiles spatial transcriptomes, immune repertoires, and microbiota from formalin-fixed paraffin-embedded (FFPE) tissues. Using Archimap, we profile the spatial localization of TCRs, BCRs, and the microbiota landscape in archived clinical tissues at single-cell resolution. Through comprehensive benchmarking, we validate Archimaps performance and fidelity. Archimap in situ assembles the immune complex and reconstructs the clonal evolution of antibodies. Together, Archimap shows the power of in situ discovery of functional immune repertoires for their antitumor immunity.

7
The Phantom of the PCR: detection and consequences of spurious UMIs in mainstream RNA sequencing

Sugino, K.; Lee, T.

2026-08-06 bioinformatics 10.64898/2026.08.01.742199 medRxiv
Top 0.1%
16.7%
Show abstract

Unique molecular identifiers (UMIs) support digital molecular counting by tagging molecules before amplification, but assume that UMIs are incorporated only during reverse transcription. Residual UMI-bearing oligonucleotides can instead reprime during preamplification PCR, creating "phantom" UMIs on genuine cDNA that inflate counts and evade standard deduplication. We model phantom generation as a two-state branching process and show that it produces a heavy-tailed reads-per-UMI distribution distinct from that of true UMIs. Using this signature, PhantomUMI detects and estimates contamination from clone-size distributions, subject to a coverage-dependent identifiability limit. Across 23 datasets spanning published studies and companion experiments, we find signatures consistent with phantom-UMI generation, including in current 10x GEM-X chemistry. Simulations show that phantoms inflate molecule counts and distort fold-changes. Model-based correction removes average count inflation but does not recover the distorted fold-changes, indicating that phantom UMIs are best prevented experimentally, as implemented in the companion Omega-seq method.

8
Learning Discrete Cell and Niche Codes from Spatial Transcriptomics Using Dual Residual Vector Quantization

Birk, S.; Merchant, A.; Vahidi, A.; Theis, F. J.; Lotfollahi, M.

2026-08-13 genomics 10.64898/2026.08.07.743490 medRxiv
Top 0.1%
15.6%
Show abstract

Spatially-resolved transcriptomics (SRT) measures gene expression at single-cell resolution while preserving each cells spatial location, enabling the joint study of cell identity and cellular niche, the recurring microenvironment that organises tissue function. Existing representation-learning methods typically capture only one of these axes at a time. We present SQUINT, a graph vector-quantized variational autoencoder (VQ-VAE) that learns two disjoint codebooks per cell from a shared architecture: a cell codebook quantising the per-cell embedding before neighbourhood aggregation, biased toward cell-intrinsic identity, and a niche codebook quantising the embedding after graph neural network (GNN) aggregation, biased toward spatial context. Both use residual vector quantization, giving a coarse-to-fine discrete-token hierarchy. SQUINT is trained with per-branch negative-binomial reconstruction objectives and three domain-motivated components that we show are crucial: a within-section cosine adjacency loss that anchors the niche codes in the spatial graph, a cross-section contrastive loss on the cell latents that aligns transcriptomically matched cells, and a decoder section covariate that absorbs batch effects. Across three datasets spanning four spatial assays (STARmap, MERFISH, CosMx, Xenium) and four tasks - niche identification, cell-type identification, cross-section integration, and spatial gene-expression imputation in held-out regions - SQUINT outperforms or is competitive with strong baselines on identification and achieves the most faithful cross-section integration. The resulting discrete vocabulary makes tissues directly consumable by transformer-style foundation models and enables one-step query-to-reference atlas mapping via code-distribution similarity, which we demonstrate on a CosMx human non-small-cell lung cancer cohort.

9
Calibration-free compression brings Evo 2 to its full million-token context on a single GPU

Patsakis, M.; Tzanakakis, A.; Georgakopoulos-Soares, I.

2026-09-01 bioinformatics 10.64898/2026.08.28.747902 medRxiv
Top 0.1%
15.3%
Show abstract

Evo 2 is the largest openly available genomic foundation model, but its forty billion parameter configuration cannot be loaded onto a single 80 GB accelerator, placing genome-scale analysis beyond most laboratories. We present TurboQuant-Bio, an open toolkit that compresses Evo 2s weights and attention cache to four bits without calibration data, and serves both through fused kernels. Compression is near-lossless across perplexity spanning the tree of life, genomic classification, splice-site prediction, gene completion and clinically relevant variant-effect prediction. It brings Evo 2 40B onto one 80 GB GPU and Evo 2 7B to its full million-token context within a 40 GB memory budget, an eightfold gain in reachable context. We further show that the released chunked-prefill path is silently incorrect, returning plausible but uncorrelated likelihoods, and derive the block-wise continuation that repairs it: a complete 580-kilobase bacterial genome is now scored in one context in 22 minutes rather than 13.7 hours.

10
A calibrated novelty flag for fungal ITS metabarcoding: choosing the error rate at which sequences are declared new

O'Brien, A.; Parada, P.

2026-08-04 bioinformatics 10.64898/2026.07.29.741524 medRxiv
Top 0.1%
15.3%
Show abstract

O_LIEnvironmental fungal surveys routinely recover internal transcribed spacer (ITS) sequences that cannot be assigned at fine taxonomic ranks, the so-called fungal "dark matter." Such sequences are set aside by thresholding a similarity or confidence score at a conventional value. Those conventions do abstain, but the error rate a threshold implies is neither stated nor selectable, and a threshold defined for one kind of score does not transfer to another. C_LIO_LIWe present a conformal novelty flag that supplies what is missing: each query receives a p-value with a distribution-free guarantee that the rate of falsely declaring a known sequence novel is bounded by a user-chosen . We evaluate it on a leave-one-genus-out benchmark built from UNITE, on alignment identity, a k-mer bootstrap consensus and two neural classifiers output probabilities, and on a soil fungal dataset. C_LIO_LIThe flag holds its nominal rate across two orders of magnitude in , so an operating point can be chosen rather than inherited: at = 0.05 it fires on 4.3% of known-genus queries and recovers 19.2% of genuinely novel genera. The cutoff holding a 5% error rate here is 64.8% identity, nowhere near the customary 97%, showing how little a threshold carries its error rate between datasets. Coverage transferred across eleven settings spanning those four scores, two amplicon regions and a sevenfold change in reference size, all within 1.1 percentage points of nominal, while detection ranged from 5.2% to 53.3%: the guarantee is on the error rate and not on power, and two of our settings are valid but uninformative. Applied to soil data the flag identifies 20.7% of amplicon sequence variants as novel at a controlled 5% error rate, 16.3% under an abundance filter. Half of those recur near-identically among GlobalFungis unnamed environmental variants while fewer than one in ten matches a named species hypothesis, a sixfold skew towards the uncatalogued against 1.9-fold for sequences the flag passes. C_LIO_LIThe flag turns an arbitrary cutoff into a decision with a stated error rate, and in doing so converts dark matter from a residue into a set of prioritizable targets for formal description. C_LI

11
High-resolution single-molecule replication profiling of the human genome

Tourancheau, A.; Rojat, V.; Ciardo, D.; Proux, F.; Lacroix, L.; Arbona, J.-M.; Audit, B.; Hyrien, O.; Le Tallec, B.

2026-07-20 genomics 10.64898/2026.07.15.738623 medRxiv
Top 0.1%
15.0%
Show abstract

Despite the considerable progress made in recent years thanks to the rise of long-read sequencing, high-resolution, single molecule (SM)-based replication profiling of large genomes has remained out of reach. Here, we present a replication labelling strategy consisting in repeated pulse-labelling of asynchronously growing cells with the thymidine analogue bromodeoxyuridine (BrdU) that greatly enhances the number of detectable replication tracks. We show that the multipulse BrdU labelling protocol, in association with ForkML, a machine-learning method translating BrdU signals in nanopore reads into oriented and positioned replication forks, enables the establishment of a SM replication map of the entire human genome at kilobase resolution.

12
Scalable single-cell isoform profiling with sequencing-by-expansion

Georgescu, C. H.; Al-Eryani, G.; Brookhart, A.; Chandrasekar, J.; Yaung, S. J.; Rogers-Peckham, M.; Freer, M.; Kartje, M. E.; Yu, H.; Khorgade, A.; Yang, C.; McGee, L.; Berg, K.; Cech, C.; Barrett, S.; Arryman, A.; Bartlett, D. A.; Slamin, A.; Low, S.; Dubinsky, D.; Cipicchio, M.; Hacohen, N.; Lehmann, T.; Lennon, N. J.; Popic, V.; Zhao, C.; Prindle, M.; Mannion, J.; Nabavi, M.; Haas, B. J.; Kokoris, M.; AlKhafaji, A. M.

2026-07-21 genetics 10.64898/2026.07.15.738809 medRxiv
Top 0.1%
15.0%
Show abstract

Single-cell RNA sequencing has transformed our understanding of cellular systems, yet the reliance on short-read sequencing restricts analysis to gene-level quantification and obscures the immense biological diversity generated by alternative splicing. While long-read sequencing technologies can capture full-length RNA and resolve transcript isoforms, current platforms remain constrained by throughput and high per-base costs, rendering them impractical for modern million-cell applications. To address this critical limitation, we developed and optimized sequencing-by-expansion (SBX) chemistry for high-throughput single-cell RNA isoform profiling. Integrated within the AXELIOS 1 sequencing platform, SBX employs a unique biochemical conversion process that transforms complementary DNA into expanded surrogate high signal-to-noise polymers called Xpandomers which are sequenced via translocation through a dense nanopore array yielding over 9.5 billion reads in a two-hour run. To leverage this unique data type for long-read single-cell RNA isoform sequencing, we developed the Consensus UMI Deduplication using Longest Length (CUDLL) algorithm, which computationally consolidates variable-length raw SBX reads into single, high-fidelity consensus reads, elevating sequence accuracy to 99.83% and maximizing per transcript read length. We demonstrate that this consensus approach successfully captures the vast isoform diversity of single-cell libraries and enables the accurate measure of differential isoform expression across distinct cell types in peripheral blood mononuclear cells. Furthermore, SBX coupled with CUDLL efficiently resolves T-cell and B-cell receptor clonotypes directly from whole-transcriptome libraries without the need for VDJ-specific target enrichment. Ultimately, this work establishes SBX and the AXELIOS 1 as a transformative platform for high-scale single-cell isoform sequencing.

13
Measuring and removing near-duplicate contamination in alignment-free SARS-CoV-2 lineage classification benchmarks

Jamhuri, M.; Irawan, A.

2026-08-18 bioinformatics 10.64898/2026.08.12.744560 medRxiv
Top 0.1%
14.9%
Show abstract

Alignment-free lineage assignment from k-mer frequency profiles is widely used for SARS-CoV-2 surveillance, and the methods that do it are ranked against each other by margins of one or two percentage points. Those rankings rest on an unchecked protocol. Public repositories hold many near-duplicate genomes, and stratified random splitting puts members of such a group on both sides of the split, so a classifier is credited for sequences it has already seen. We propose quantised profile hashing, which finds near duplicates in k-mer feature space by rounding each frequency vector and hashing it. No sequence is compared with any other, so one pass over the feature matrix suffices and no similarity threshold has to be chosen. Rounding is also what makes the groups well defined, and they are then kept whole across the training, validation and test sets. On 255,611 genomes from seven Pango lineages, random splitting leaves 5.09% of test sequences with a near duplicate in training, on a benchmark ranked by margins of one or two points. Ten update rules were trained twice, identically except for the partition. The contaminated benchmark separates one rule from the leader at 0.05; the clean one separates none. The two orderings are uncorrelated, Kendall{tau} = +0.022, with rules moving 3.2 positions on average and the leader of one benchmark ranking eighth on the other. A ranking obtained under contamination therefore says nothing about the ranking without it, and the quantity worth reporting beside a score is the leakage rate of the split.

14
Genomic foundation model embeddings encode higher-order viral genome architecture beyond sequence composition: a benchmark of Evo 2

Amgarten, D.; Schinaid, A.; de Mello Malta, F.; Marra, A. R.; Rebello Pinho, J. R.

2026-07-16 bioinformatics 10.64898/2026.07.14.738542 medRxiv
Top 0.1%
14.7%
Show abstract

Genomic foundation models such as Evo 2 are increasingly applied to microbial genomics, yet how well their representations capture viral genome organisation, and how reliably they generate viral sequence, remain poorly characterised. We present a reproducible benchmark of Evo 2 on viral genomes. Using a pre-registered RefSeq viral corpus (19,429 genomes, organised by Baltimore class and host domain), we evaluated three axes: linear probes decoding Baltimore class, host domain and viral family from mean-pooled embeddings; ridge-regression probes recovering genomic features, including higher-order architectural properties such as gene density, coding fraction and gene overlap; and generative completion of fragmented genomes, scored on a leakage-safe set of eukaryote-infecting viruses (excluded from Evo 2s training corpus by design) against a bacteriophage comparator. All probes used cross-validation with sequence-identity-clustered folds, benchmarked against both a GC-and-length control and a 6-mer composition representation. From its optimal intermediate layer, the 20B embedding classified Baltimore class at 0.96 accuracy and host domain at 0.99, exceeding both baselines; for viral family, however, 6-mer composition (0.89) matched the embedding (0.91. Most informatively, the embedding decoded coding fraction, gene density and gene overlap (R{superscript 2} = 0.61, 0.77 and 0.64) far beyond 6-mer composition (0.10, 0.38 and 0.27), evidencing genuine encoding of genome architecture rather than nucleotide composition (p < 0.001). Performance scaled with model size. In generation, perplexity was lower for bacteriophages (1.18 bits/nt) than for held-out eukaryotic viruses (1.80). Evo 2 encodes functional viral genome architecture beyond composition, while taxonomic and generative behaviour partly reflect composition and training exposure.

15
Evaluating cell type annotations in single-cell omics in the absence of ground truth

Garnica, J.; Andreatta, M.; Carmona, S. J.

2026-06-12 bioinformatics 10.64898/2026.06.10.731285 medRxiv
Top 0.1%
13.4%
Show abstract

Accurate cell type annotation is essential for single-cell transcriptomics, directly shaping downstream analyses and biological interpretations. Yet, objective evaluation of annotation quality remains a major challenge. Here, we argue that a cell type or cell state label has practical utility only if it captures a molecular pattern that is reproducible across biological replicates. Based on this principle, we introduce inter-sample consistency (ISC), a quantitative framework to assess annotation quality in single-cell RNA-seq datasets. Unlike existing cluster validation approaches, ISC distinguishes annotations that generalize across samples and individuals from those driven by technical or unwanted variation, thereby providing principled criteria for annotation quality and transferability. When applied to published single-cell atlases, ISC reveals widespread reproducibility gaps and provides actionable guidance for repairing inconsistent annotations. Notably, ISC enables benchmarking of automated cell type annotation tools even when ground-truth labels are unavailable, providing interpretable metrics to guide their development and evaluation. Implemented as the scTypeEval Bioconductor package, this framework offers a broadly applicable resource for evaluating and improving cell type annotations in single-cell RNA-seq experiments.

16
Cross-architecture ensembling of DNA foundation models improves the precision and stability of chimera detection in long-read metagenomic bins

MinSeo, K.; Jae-Ho, S.

2026-07-07 bioinformatics 10.64898/2026.07.02.735979 medRxiv
Top 0.1%
13.3%
Show abstract

Motivation: Chimeric metagenome-assembled genomes (MAGs) that pool DNA from multiple organisms contaminate downstream analyses. Marker-gene tools such as CheckM2 miss low-level chimerism, and DNA foundation models have been proposed as a sequence-composition alternative, but whether large autoregressive models (Evo2, 7B parameters) outperform smaller contrastive models (DNABERT-S, 117M) has not been rigorously tested.

17
Spatial multi omics enables single cell transcriptome metabolome inference

shen, x.; ZHANG, X.-Y.

2026-08-12 bioinformatics 10.64898/2026.08.06.743252 medRxiv
Top 0.1%
13.0%
Show abstract

Joint single-cell transcriptomic-metabolomic profiling remains technically intractable. Here we present CHIMERA (Cell-level Hybrid Inference of Metabolome Embedded on RNA Atlas), a data-driven framework that learns transcriptome-to-metabolome mappings from spatially paired multi-omics data and transfers them to unpaired scRNA-seq. CHIMERA generates quantitative, database-independent single-cell metabolite abundances and, by pairing them with the measured transcriptome of the same cells, enables joint co-embedding of genes and metabolites for the discovery of differential metabolites and co-regulated gene-metabolite modules. Using 10x Visium paired with MALDI-MSI from murine liver sections and a matched scRNA-seq reference, CHIMERA achieves a per-metabolite median Pearson r = 0.285 with positive cross-section generalization. On an independent Liver Cell Atlas Western-diet cohort, CHIMERA recovers metabolic reprogramming that recapitulate published non-alcoholic fatty liver disease pathophysiology. Applied to a Rarres2 (chemerin) knock-down hepatocellular carcinoma model, CHIMERA uncovers metabolic heterogeneity among tumour-associated macrophages, resolving four metabolic subclusters (MC-0 to MC-3); Rarres2 appears to drive macrophage polarization from an LAM-like MC-3 state toward Spp1+ like MC-0/MC-2 by modulating a co-regulated gene-metabolite module--a dual-omics phenotype undetectable by either modality alone. CHIMERA is the first data-driven framework for quantitative single-cell metabolome inference, opening joint transcriptomic- metabolomic analyses inaccessible to either experimental or knowledge-based computational approaches.

18
WattmaMod enables high-resolution and extensible RNA modification profiling for nanopore direct RNA sequencing

Han, R.; Yu, B.; Xinghui, S.; Xiao, L.; Junhai, Q.; Ting, Y.; Xin, G.

2026-07-02 bioinformatics 10.64898/2026.07.02.735990 medRxiv
Top 0.1%
12.9%
Show abstract

Nanopore direct RNA sequencing enables direct profiling of RNA modifications on native transcripts, but accurate multi-modification detection remains limited by non-stationary signals and heterogeneity across chemistries. Here, we develop WattmaMod, a deep learning framework for multi-modification detection from nanopore direct RNA sequencing data. It combines self-supervised pretraining, supervised contrastive fine-tuning, and low-label incremental adaptation to improve representation learning and support efficient extension to low-resource modification types. The framework further incorporates wavelet-guided multi-scale encoding and dynamic cross-attention fusion to model raw signals and event-level features. Results show that WattmaMod achieves robust detection of multiple RNA modifications, including m6A, m5C, m1A, A-to-I, m7G, hm5C, m1{Psi}, f5C, ac4C, m5U and {Psi}. It also extends efficiently to low-resource modification types with minimal labeled data, generalizes across sequencing chemistries and species, and predicts potential higher-order local organization among distinct RNA modifications. WattmaMod thus provides a scalable framework for high-resolution epitranscriptome profiling and expands RNA modification analysis beyond single-site prediction to coordinated multi-modification characterization.

19
HuMMANet: A Harmonized Cross-Study Resource for Integrative Analysis of Human Gut Microbiome Metabolome Associations

Verma, S.; Arora, N.; Ajay, C. P.; Singh, P.; Mallick, H.; Ghosh, T. S.

2026-08-25 bioinformatics 10.64898/2026.08.24.746727 medRxiv
Top 0.1%
12.8%
Show abstract

Deciphering gut microbiome to host metabolome interaction is critical for understanding how microbial communities generate bioactive signals that shape host physiology and disease. Progress, however, has been hindered by inconsistent metabolite annotations, poor interoperability across studies, and the absence of integrated resources placing microbiome-derived metabolites within their functional, microbial, physiological, and clinical context. Here we present HuMMANet (Human Microbiome Metabolome Annotation Network), a harmonized resource integrating 46 paired gut microbiome metabolome studies (59 study-units; 14,405 samples; 13 disease categories plus a healthy/control reference category) with a scalable metabolite-harmonization framework. HuMMANet resolves heterogeneous annotations through a multi-stage workflow spanning RefMet, HMDB, PubChem, Metabolomics Workbench, SMPDB, MiMeDB 2.0, GNPS/microbeMASST, DrugBank, and DrugCentral, yielding a reference atlas of 54,914 unique metabolites, annotated with standardized chemical identifiers, biochemical pathways, microbial producer associations, physiological distributions, disease links, and structural relationships to approved therapeutics, a unified reference framework for microbiome metabolome research. Applying HuMMANet to a multi-cohort integration of adult serum and fecal metabolomes, we identified 519 serum and 322 fecal metabolites reproducibly associated with gut microbial community composition (PERMANOVA, P < 0.05 in at least 50% of studies in which detected), enriched for specific biomolecular classes and pathways. Cross-referencing these against Health Associated Core Keystone (HACK) taxa revealed 58 serum and 25 fecal metabolites (HACK positive) whose taxon-level associations tracked positively with the taxon specific HACK indices. These reproducible metabolomic signatures of microbiome health included indole3propionic acid, a gut barrier-protective microbial tryptophan metabolite, and 3phenylpropionate. Drug similarity annotation within HuMMANet linked 16 of this serum and 13 fecal HACK positive metabolites to therapeutics used in neurological, inflammatory, and vascular disease. Conversely, 38 serum and 65 fecal metabolites, including imidazole propionate and long-chain acylcarnitines such as ACar 18:0, showed HACK negative signatures previously associated with dysbiosis-linked disease. GNPS/microbeMASST and MiMeDB 2.0 annotations further traced subsets of these metabolites to putative bacterial producers. HuMMANet thus provides a standardized framework for reproducible microbiome metabolome integration, enabling cross study discovery and translational prioritization of conserved microbiome derived metabolic signatures across human populations and disease states.

20
Quantitative profiling of intrinsic dCas9-DNA recognition reveals key determinants of guide RNA performance

Zhu, W.; Tian, M.; Duan, Y.; Reisman, S. J.; Miller, S. E.; Corden, E.; ter Weele, M.; Song, L.; Blount, J.; Safi, A.; Schreiber, J.; Gersbach, C. A.; Crawford, G. E.; Gordan, R.

2026-08-10 genomics 10.64898/2026.08.10.743836 medRxiv
Top 0.1%
12.7%
Show abstract

CRISPR technologies based on nuclease-deactivated Cas9 (dCas9) rely on programmable DNA binding rather than DNA cleavage, yet the intrinsic DNA-recognition properties that govern optimal guide RNA (gRNA) performance remain poorly understood. Existing approaches either measure genomic occupancy in cells or infer dCas9 behavior from cleavage-based Cas9 datasets, despite DNA binding being substantially more permissive than DNA cleavage. Here we introduce TANGO (Targeted Array-based Nucleic acid-Guided Occupancy), a high-density DNA-array platform that quantitatively profiles intrinsic dCas9:gRNA binding across tens of thousands of DNA targets in a cell-free system. TANGO captures established features of dCas9 target recognition, while providing substantially greater sensitivity than prior assays. Comparison with ChIP-seq data demonstrates that intrinsic DNA-binding specificity is a major driver of genomic occupancy and reveals that chromatin accessibility modulates the intrinsic binding affinity required for dCas9 recruitment. Across CRISPRi/a guides, TANGO identifies multiple independent biochemical determinants of guide performance--including on-target affinity, mismatch tolerance, and ribonucleoprotein assembly--and flags problematic and highly promiscuous guides overlooked by current specificity metrics. Unexpectedly, some guides retain substantial guide-directed DNA binding even in the absence of a protospacer-adjacent motif (PAM), revealing an additional dimension of dCas9 specificity. Together, these results establish intrinsic DNA recognition as a quantitative and experimentally accessible determinant of dCas9 function, providing a framework for improving guide selection and enhancing the precision of CRISPR technologies.